Skip to content

fix(forecast): choose each forecast source from measured errors - #1494

Merged
frahlg merged 3 commits into
masterfrom
1490-source-choice
Oct 1, 2026
Merged

frahlg merged 3 commits into
masterfrom
1490-source-choice

Conversation

@frahlg

@frahlg frahlg commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Problem

Core picks each forecast signal's source from the worker's quality label alone:

  • PV uses Energyplan even in cold_start.
  • Load rejects Energyplan in cold_start and falls back to the legacy twin.

The archive already scores both sources against measurements, but the choice ignores it. Home box, 22 Sep–1 Oct, ~1,600 issues (bias / MAE):

Source Night load Day load Day PV
Used by the plan +1,979 / 2,006 W +890 / 1,381 W −406 / 1,405 W
Energyplan +38 / 381 W −256 / 786 W −749 / 1,790 W
Legacy +2,402 / 2,402 W +1,020 / 1,503 W −166 / 1,095 W

A closed-loop replay of 24 Sep–1 Oct with Energyplan load cut the week's grid cost by 16 SEK (24 SEK with the Balanced margin from #1486), about half the gap to a perfect forecast.

Change

  • chooseForecastSources compares energyplan and legacy_shadow errors from the same issues, in the current cohort, over the last 7 days. forecasting.PoolFrozenSeries pools the lead buckets and counts distinct scored hours and days. It uses pv_daylight for PV and load for load.
  • With at least 48 scored hours over at least 3 days, and one source at least 10 % better by MAE, that source wins for the signal, even if Energyplan load is still cold_start. Otherwise today's quality rule decides. Legacy PV wins only where legacy has weather for the interval; elsewhere Energyplan PV is used as before.
  • Each slot's planning margin comes from the errors its own sources made on earlier issues: the energyplan series, the legacy_shadow series, or a mix that forecasting.ComposeFrozenSeries builds from both on the same issues. A source switch therefore keeps its history, and a band never describes a source the plan stopped using. Archived champion bands are unchanged.
  • Core logs forecast source chosen with hours, days and both MAEs whenever a choice changes. Each issue already records the source of every point.
  • forecastPipelinePolicy moves to energyplan-primary-v3, so bands never mix errors from the old rule. Forecast error bands reset on every Core update #1489 already starts one new cohort in this release, so this costs no extra reset.
  • docs/architecture.md and docs/energyplan-contract.md state the new rule.

What a site sees

For the first three days after the update the quality rule decides, as today. Then each signal follows its measured errors. On the home box that would mean Energyplan load and legacy PV.

When cold Energyplan load wins on evidence, the load margin uses its model range until empirical bands exist (#1493). That is the honest range of the source in use.

Cost: with a full 30-day archive (34,560 errors), the added margin calibration takes 14 ms and the choice 2 ms on a Mac, next to 17 ms for the existing calibrator. On a Pi 4 that is roughly a tenth of a second per replan.

Tests

  • TestForecastSourceChoiceFollowsMeasuredErrors: enough evidence picks per signal; one day, a small gap, another cohort or week-old evidence keep the quality rule; reported MAEs match.
  • TestForecastSourceChoiceSteersResolve: evidence puts cold Energyplan load and legacy PV into the plan and records both sources; a clearly better legacy load beats a learned Energyplan load.
  • TestForecastRiskUsesTheSlotsOwnSources: after a measured choice falls back to the quality rule, and after legacy load wins, the margin comes from the slot's own sources, not from champion errors of the earlier source.
  • TestPoolFrozenSeriesCountsEachTargetHourOnce, TestComposeFrozenSeriesTakesEachSignalFromItsSeries in internal/forecasting.
  • Each test fails with its rule disabled. make verify clean.

Review

A local Codex review found three problems, all fixed here: calibration went back to mixed champion errors when the choice returned to the quality rule; a legacy PV choice blocked Energyplan PV where legacy had no weather; the evidence gate counted hours in one lead bucket only. A second Codex pass found nothing further.

Fixes #1490. One of four forecast fixes found while evaluating #1482, with #1492, #1493 and #1495. Rebased onto #1493.

🤖 Generated with Claude Code

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-01T17:27:49.218444Z 73db64d PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 73db64dbbd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines 453 to +454
calibrator := forecasting.NewCalibrator(history, site.Revision, origin.UnixMilli())
choice := chooseForecastSources(history, site.Revision, origin.UnixMilli())

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reset calibration when the selected source changes

When the evidence threshold switches a signal from the quality-rule source to the measured winner, site.Revision remains unchanged, so this calibrator continues using champion residuals produced by the previously selected source. Once the seven-day empirical threshold is reached, Risk can therefore apply a mixed or legacy-source band to the newly selected Energyplan forecast (and vice versa), bypassing that source's model bounds and producing an incorrect planning margin. Key calibration by the recorded signal source or start a new calibration cohort whenever the choice changes.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 15d4045: calibration now keeps only champion errors whose recorded sources match the current choice. A new cohort per switch would also reset the choice evidence, which is filtered by cohort, so the choice could flip back and forth. After a switch the bands start cold until the new choice has its own errors. Test: TestForecastSourceChoiceCalibratesOnlyTheChosenSources.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: the filter in 15d4045 was replaced in c9dc72b. Each slot's margin now calibrates on the errors its own sources made (the energyplan or legacy_shadow series, or a mix composed from both on the same issues), so a switch in either direction keeps its history. Test: TestForecastRiskUsesTheSlotsOwnSources.

frahlg and others added 3 commits October 1, 2026 20:10
Core picked PV and load sources from the worker's quality label alone.
On the home box that put cold Energyplan solar and a legacy load that
overshot by 2 kW at night into the plan, while the archive already
showed the other source was better for each signal.

Compare Energyplan and legacy_shadow errors from the same issues over
the last week. With 48 scored hours over three days and a 10% gap, use
the better source for that signal; otherwise keep the quality rule.
Log each change of choice. Bump the pipeline policy so bands do not
mix errors from the old rule.

Fixes #1490

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Fredrik Ahlgren <fredrik@sourceful-labs.com>
After the choice moves a signal to another source, champion errors
from the old source stayed in the calibration window for up to 256
hours, so the margin described a source the plan no longer used.
Keep only champion errors whose recorded sources match the current
choice. Until the new choice has its own errors, bands start cold.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Fredrik Ahlgren <fredrik@sourceful-labs.com>
Three findings from a local Codex review:

- Calibration filtered champion errors only while a measured choice
  held. When the choice fell back to the quality rule, old errors of
  another source set the margin again. Each slot's margin now comes
  from the errors its own sources made on earlier issues: Energyplan,
  legacy, or a mix composed from both series on the same issues. A
  switch keeps its history instead of starting cold.
- A measured legacy PV choice also blocked Energyplan PV where legacy
  had no weather. Keep legacy PV only where it is known.
- The evidence gate counted hours in the largest lead bucket. Count
  distinct scored hours and days across all buckets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Fredrik Ahlgren <fredrik@sourceful-labs.com>
@frahlg
frahlg force-pushed the 1490-source-choice branch from 15d4045 to c9dc72b Compare October 1, 2026 18:10
@frahlg
frahlg merged commit 1757ce0 into master Oct 1, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Forecast source choice ignores measured accuracy

1 participant